Computer Methods and Programs in Biomedicine
○ Elsevier BV
Preprints posted in the last 7 days, ranked by how well they match Computer Methods and Programs in Biomedicine's content profile, based on 28 papers previously published here. The average preprint has a 0.04% match score for this journal, so anything above that is already an above-average fit.
Oyarzun, R.; Hernandez, P.
Show abstract
Background. Whether predictors of intraoperative hypotension (IOH) add information beyond the mean arterial pressure (MAP) already displayed on the monitor is contested: selection bias in common evaluation designs inflates apparent performance, and the field has called for comparisons against simple MAP-based references under bias-resistant protocols. Existing predictors also depend on proprietary waveform analysis or pulse-contour monitors, restricting both deployment and external validation. Methods. Using 807 non-cardiac surgery patients from the open VitalDB database, we derived an additive gradient boosting model (one split per tree: a learned shape function per variable, no interactions) from three variables computable from an arterial line alone: current MAP, its drift from the patient's own 20-minute baseline, and the growth of its rolling variance (critical slowing down). Evaluation used patient-level 5-fold cross-validation under a strict protocol - exclusion of the 65-75 mmHg grey zone and of all samples already hypotensive at prediction time - with MAP alone (same learner class) as comparator. The frozen model was then validated, without any refitting, on an independent cohort from another continent (MOVER, University of California Irvine) following a pre-registered plan sealed before external data access. Results. In development the pressure-only model reached AUROC 0.907 vs. 0.884 for MAP alone (Delta AUROC +0.023, 95% CI +0.017 to +0.029) at 5 min, with +0.031 and +0.032 at 10 and 15 min, and good calibration (Brier skill +0.418 vs. prevalence). In external validation on 3,069 patients (442,194 samples, 1-minute charting, event prevalence 5.8%), the advantage not only transferred but was larger than in development: AUROC 0.696 vs. 0.638, Delta AUROC +0.058 (95% CI +0.051 to +0.064), meeting both pre-registered gates. Discrimination transferred; calibration did not (external Brier skill -0.014), requiring local recalibration. In the unrestricted scenario, where samples already at threshold are retained, the advantage collapsed (+0.007), reproducing the selection effect this paper documents. A secondary model adding pulse-contour cardiac output and stroke volume variation improved development discrimination further (Delta AUROC +0.035) but could be externally validated in only 39 patients, because those signals are rarely recorded. Conclusions. The dynamics of arterial pressure itself - drift from a patient-specific baseline and variance growth - carry predictive information beyond its current value, in a fully interpretable additive model that requires only an arterial line, no waveform access and no proprietary hardware. The advantage is confirmed in a pre-registered frozen-model external validation of over three thousand patients, and is largest at coarse recording cadence, where instantaneous pressure is least informative.
Oyarzun-Silva, R. A.; Hernandez-Hernandez, P.; Fernandez-Vaquero, M. A.; De Luis-Cabezon, N.
Show abstract
Background. Videolaryngoscopy still requires adjuncts or hyperangulated rescue in a clinically important minority, and bedside screening discriminates modestly. Point-of-care ultrasound (POCUS) of the anterior airway is a promising alternative, but existing prediction models are opaque or assume a pre-specified functional form. We developed and internally validated a parsimonious, fully disclosed POCUS risk equation whose form is recovered from data and whose structural properties are machine-checked by formal proof - to our knowledge the first formally verified clinical risk predictor - following TRIPOD+AI 2024. Methods. In a prospective single-centre, single-operator cohort of 259 adults undergoing elective videolaryngoscopy (no-Easy airway 68/259, 26.3%), Sequentially Thresholded Least Squares with bootstrap stability selection (B=300) screened a 71-term library of nine POCUS features and retained a seven-term logistic equation; a two-term bootstrap-stable model was pre-specified as robustness analysis. Internal validation used 5x10 repeated cross-validation plus temporal and device hold-outs, with pre-specified overfitting and optimism assessments. Five behavioural properties of the deployed equation were machine-checked in Lean 4. Results. Two interactions met the |c|/sigma_c>2 stability criterion: skin-to-epiglottis x skin-to-hyoid-bone distance and tongue volume x sagittal tongue area. The seven-term equation reached a 5x10 cross-validated C-statistic of 0.966 (optimism-corrected 0.968) and held across temporal and device hold-outs (0.94-0.97). Calibration-in-the-large matched prevalence, with cross-validated slope 0.90 attenuating to 0.625 out-of-time; standard recalibration restored 0.92 without loss of discrimination. The pre-specified two-term robustness model reproduced this performance (C-statistic 0.964-0.968; events-per-parameter 34; shrinkage 0.99), confirming the result is not an artefact of the screening stage. Net benefit over a clinical baseline was positive across 10-50% thresholds. All five Lean 4 theorems compiled without sorry. Conclusions. A sparse, formally verified POCUS equation predicts difficult videolaryngoscopy with high internally validated discrimination and quantified, modest overfitting. Because the equation was developed in a single-operator cohort and its inputs are operator-dependent, external validation requires prior harmonisation of the measurement protocol and operator credentialing.
Ekambarapu, L.; Pendyal, A.; Lin, A.; Alwakeel, M.; Rajaratnam, A.
Show abstract
Background: Unstructured biomedical data, such as echocardiography reports, are rich in information but time consuming to analyze at scale. Rule-based, regular expression-driven terminology mapping can only extract individual variables while large language models (LLMs) offer scalable and clinically meaningful interpretations of heterogeneous disease processes. Right ventricular dysfunction (RVD) is an example of a multifactorial disease state in which key structural and physiologic features are captured both narratively and in structured fields, making it an ideal test case for evaluating whether LLMs can recover complex phenotypes that rules based methods routinely miss. Purpose: To compare an LLM-based extraction method to a conventional rules-based schema for identifying and phenotyping echocardiographic features associated with RVD in a large TTE dataset. Methods: MIMIC-III NOTE2NUM echocardiography reports (n = 45,794) were analyzed using GPT-4o-based LLM extraction deployed within a secure health system enclave and were benchmarked against echocardiographic measurements defined in the MIMIC-III dictionary schema. In MIMIC-III, PH was recorded qualitatively (mild/moderate/severe) based on tricuspid regurgitant (TR) jet velocity and then re-coded as present vs. absent. LLM based extraction defined RVD as (1) RV structural abnormality (>= 1 of hypertrophy, dilation, or wall hypo-/akinesis) or (2) RV pressure/volume overload (>= 2 of the following: estimated right atrial pressure > 8 mmHg, TR jet velocity > 2.8 m/s, fractional area change < 35%, tricuspid annular planar systolic excursion < 17 mm, S' < 9.5 cm/s, or E/e' > 14), with PH defined as estimated pulmonary artery systolic pressure > 35 mmHg or qualitative documentation of PH. Results: LLM extraction identified PH in 15,394 (33.6%), RV pressure/volume overload in 14,449 (31.6%), and RV structural abnormalities in 11,955 (26.1%). Co-occurrence was common: overload + structural changes in 9,380 (20.5%), overload + PH in 9,756 (21.3%), structural changes + PH in 6,183 (13.5%), and all three in 5,620 (12.3%). Using the MIMIC-III dictionary schema, PH prevalence was similar (15,371; 33.6%), but RV overload fields were captured less often (pressure overload 1,357 [3.0%], volume overload 1,128 [2.5%], pressure + volume overload 1,093 [2.4%]; any overload field 3,578 [7.8%]), and RV pressure/volume overload with PH was identified in only 731 (1.6%). Conclusions: LLM-based extraction outperforms rules-based schemas for identifying complex disease states not defined by any single variable. By synthesizing multifactorial signals, LLMs can phenotype RVD with higher fidelity and support population-level assessment. Further validation using multimodality imaging, invasive hemodynamics, and clinical outcome data is needed.
Okundaye, D. O.; Isiekwene, C. C.
Show abstract
Acute kidney injury (AKI) is a frequent complication within intensive care units, with its sudden onset often missed. This is especially important because a timely window for intervention is required as delayed detection leads to progressively worse outcomes. Existing machine learning and deep learning models have contributed to closing this gap, but their complexity, requiring hundreds to thousands of features, and lack of generalisation pose a limitation that prevents them from being integrated into clinical workflows across different electronic health-record ecosystems. This study presents a 37-feature XGBoost model trained on the MIMIC-IV dataset with 5.4% positive cases, with hyperparameters optimised via Optuna and probabilities calibrated using isotonic regression, designed for transportability across clinical settings. Validation was conducted internally using a temporal patient-level split simulating prospective deployment, training on 2008-2016 data and testing on 2017-2022 data"External validation was performed on the eICU Collaborative Research Database, a multi-centre dataset spanning 208 US hospitals, using the trained model without retraining. SHAP TreeExplainer was used to provide feature-level explainability for individual predictions. Internal testing yielded an AUROC score of 0.794 for predicting AKI onset within a 12-24 hour window. External validation produced a 0.750 AUROC without retraining. Equitable discrimination was observed across gender, age, chronic kidney disease presence, race, and AKI stages on both datasets, with a 95% internal CI of 0.789-0.799 confirming the model's estimate stability. These results suggest that clinically useful prediction systems are achievable with substantially fewer features than current models require.
Kohler, S.; Meyer-Eschenbach, F.; Michelena, X.; Marschollek, M.; Eils, R.
Show abstract
The openEHR standard provides an open, vendor-neutral architecture for clinical data repositories (CDRs), yet its real-world deployment has not been systematically documented. We conducted a dual-perspective survey combining a vendor survey of openEHR CDR providers with a community survey of openEHR practitioners. Eleven vendor organisations reported deployments across 22 countries and over 100 institutions and health regions. A complementary community survey (n=29, 17 countries) provided context on regulatory environments, adoption drivers, and barriers. Combined, the surveys cover 28 countries, 26 of them with a reported openEHR CDR deployment. Three findings emerge: openEHR has achieved national-scale presence through two distinct channels. Through vendor-market convergence, openEHR-based systems cover the majority of regional health authorities without a national mandate, including 19 of 21 Swedish regions, 3 of 4 Norwegian health regions, and 16 of 21 Finnish wellbeing services counties. Through national health record adoption, governments have built or procured national systems on openEHR as their technical foundation, including Ireland, Malta, Greece, Jamaica and Slovenia. Across Europe, this constitutes an openEHR-based interoperability infrastructure already in place across multiple EU member states. We identified no country in which openEHR is named in binding national regulation, creating structural fragility and an unrealised opportunity for alignment with the European Health Data Space (EHDS). Second, 61% of deployments serve primary use only, and 12% support both primary and secondary use. Third, lack of openEHR-specific knowledge is the most consistent adoption barrier across all geographies and deployment tiers. Adoption is driven by practitioner need and innovation, not by regulatory mandate.
Greendyk, J. D.; Allen, W. E.; Hossain, A.; Trichas, Z.
Show abstract
Background: Percutaneous mechanical circulatory support (pMCS) is increasingly used in critically ill patients, yet its value in relation to cost and outcomes remains unclear. We evaluated national variation in utilization, outcomes, and cost, and introduced a value of care framework integrating risk-adjusted outcomes and expenditures. Methods: We performed a retrospective cohort study using the National Inpatient Sample to identify non-elective hospitalizations of critically ill patients undergoing intra-aortic balloon pump (IABP) or percutaneous left ventricular assist device (pLVAD) placement using ICD-10 codes. Multivariable logistic regression and generalized linear models were used to estimate expected outcomes and costs. Observed-to-expected (O/E) ratios were calculated, and a value index was derived to compare procedural strategies. Results: A total of 57,910 weighted hospitalizations were included (IABP 78%, pLVAD 22%). In-hospital mortality exceeded 30% across regions. Significant regional variation was observed, with the West demonstrating the highest costs and the Midwest the lowest (p<0.001). Mean hospital charges were higher for pLVAD compared with IABP ($403,731 vs $320,769). Both strategies achieved outcomes better than expected after risk adjustment (O/E 0.92); however, costs were higher than expected for both, with greater relative cost inflation observed for IABP (O/E 1.41) and higher absolute costs for pLVAD. In value-of-care analysis, IABP was associated with lower cost and comparable outcomes, while pLVAD demonstrated higher cost without proportional outcome improvement. Conclusion: Substantial variation exists in the cost, outcomes, and value of pMCS strategies. While both IABP and pLVAD achieve favorable risk-adjusted outcomes, pLVAD is associated with higher costs without commensurate clinical benefit.
ye, y.; Zeng, Z.; Tian, X.; Yuan, Z.; Wang, J.; Zhu, Y.
Show abstract
Artificial intelligence applied to routine electrocardiograms (ECGs) has largely focused on detecting existing disease or predicting individual cardiovascular outcomes. Whether ECGs can support prediction of multiple future diseases across organ systems remains unclear. We developed ECG-RISK, a multitask survival model for 67 incident three-character ICD-10 endpoints using ECG waveforms, demographic characteristics and routinely collected laboratory data from 86,673 MIMIC-IV patients. Discrimination was highest for heart, brain, kidney and lung endpoints, with organ-level C-indices ranging from 0.796 to 0.825, whereas liver and pancreatic endpoints showed lower discrimination. The ECG-only model achieved strong discrimination across most endpoints, whereas the incremental improvement gained by incorporating ECG and laboratory inputs beyond demographic information varied substantially across endpoints. Across the nine exploratory aggregated outcomes, Kaplan Meier curves showed clear separation among model-score tertiles. Discrimination was highest for dementia (C-index, 0.891) and heart failure (C-index, 0.857). These findings support the feasibility of ECG-based longitudinal risk prediction across multiple diseases. External validation and competing-risk analyses are required to assess generalisability and clinical utility.
Song, Q.; Ni, C.; Liu, W.; Li, Y.; Malin, B. A.; Yin, Z.
Show abstract
Automatic coding from clinical notes has been studied extensively for International Classification of Diseases (ICD) codes, yet broad Current Procedural Terminology (CPT) and Healthcare Common Procedure Coding System (HCPCS) recommendation remains comparatively underexplored. Existing studies often focus on one specialty, a limited code vocabulary, or a single model family, leaving it unclear how different artificial intelligence (AI) paradigms perform under a common, clinically meaningful evaluation. We formulate CPT and HCPCS coding as an AI-assisted recommendation task in which a physician or professional coder reviews a short, ranked list of candidate codes supported by the clinical note. Using operative notes from Vanderbilt University Medical Center (VUMC) and discharge summaries from Medical Information Mart for Intensive Care IV (MIMIC-IV), we compare lexical retrieval, Clinical-Longformer, GPT-5.6-Sol, MedGemma-27B, and an inspectable agentic-style retrieve-and-verify system under a controlled review budget. Micro-averaged recall within a fixed number of recommendations measures whether reference codes reach the reviewable list; micro-F1 is reported only where reference labels are sufficiently complete. Zero-shot GPT-5.6-Sol achieves the highest recall within five and ten candidates: 0.717 and 0.800 on VUMC and lower-bound values of 0.689 and 0.738 on MIMIC-IV. The retrieve-and-verify system reaches 0.695 and 0.784 on VUMC and lower-bound values of 0.575 and 0.657 on MIMIC-IV, with a candidate-linked evidence window attached to each retained recommendation. Diagnostic analyses reveal distinct failure sources, including output-length underfilling, confusion among closely related codes, out-of-knowledge-base generation, and incomplete evidence support. These findings establish a systematic evaluation framework for procedure-code recommendation and identify practical requirements for future systems that are accurate, review-efficient, and grounded in clinical evidence.
De Luca, S.; Fava, C.; Rizzo, G.; Visconti, A.; Berchialla, P.
Show abstract
Background. Patient stratification from multi-omics and clinical data is essential for uncovering disease heterogeneity and moving toward more personalized treatment strategies. However, integrating heterogeneous data layers while identifying robust patient strata remains challenging. Methods. We introduce Reduced Fusion of Multi-Omics Stratification (RedFuMOS), a novel three-step approach for patient stratification based on mixed-type multi-omics data. RedFuMOS extends Similarity Network Fusion to accommodate mixed-type data layers and layer-specific similarity measures for data integration, includes a dimensionality reduction step to mitigate the curse of dimensionality, and performs patient stratification using density-based hierarchical clustering with HDBSCAN. It also implemented an automated optimization procedure to identify the best set of hyperparameters, minimizing the need for manual tuning. Results. RedFuMOS outperformed six state-of-the-art tools for multi-omics patient stratification in a comprehensive simulated benchmarking study, which also confirmed that, although computationally expensive, the dimensionality reduction step is crucial for achieving good stratification performance. Additionally, RedFuMOS identified two clinically relevant patient strata in a small real-world cohort of patients with Philadelphia chromosome-positive chronic myeloid leukaemia. Conclusion. RedFuMOS provides a flexible framework for integrating heterogeneous multi-omics and clinical data. RedFuMOS is available as an R package at http://github.com/delucasara/RedFuMOS.
Jaber, A.; Hughes, L.; Cameron, A. C.; Quinn, T. J.
Show abstract
Background: Systematic reviews of clinical prediction models increasingly include studies using artificial intelligence (AI) and machine learning (ML) methods alongside traditional multivariable regression approaches. A previously published Excel tool enabled standardised data extraction using the CHARMS checklist and risk of bias assessment using PROBAST. The recent publication of the PROBAST+AI framework, which distinguishes the assessment of model development quality from the assessment of model evaluation risk of bias and assesses applicability in both parts, necessitates an updated digital instrument applicable across prediction modelling methods. Methods: We updated an open-access Excel tool to incorporate the full PROBAST+AI framework. The updated template incorporates structural separation between assessment of model development quality and model evaluation risk of bias, with applicability assessed in both parts. It also incorporates updated signalling questions, including those addressing methodological issues particularly relevant to AI/ML, and automates the generation of summary tables and graphical displays. Results: The updated tool (CHARMS & PROBAST+AI Template) contains 11 worksheets and supports data extraction and appraisal for up to 30 prediction models. Dedicated, linked worksheets enable separate assessment of model development and model evaluation, with Domain 4 distinguishing among Apparent, Internal, and External evaluation settings. Key updates include dedicated assessments for predictor pre-processing, class imbalance handling and recalibration, data leakage prevention, and replication of the full model development pipeline within resampling procedures. Automated sheets dynamically format tables and summary charts covering PROBAST+AI parts. Conclusions: The CHARMS & PROBAST+AI Excel template provides a standardised, user-friendly, and rigorous digital framework for systematic reviewers appraising traditional statistical and AI-driven clinical prediction models.
Li, Z.; Fujisawa, T.; Skadberg, O.; Fineran, P.; Thurston, A. J.; Tew, Y. Y.; Aakre, K. M.; Mills, N. L.; Wereski, R.; the POC-ET Investigators,
Show abstract
Background: High-sensitivity cardiac troponin (hs-cTn) assays enable safe early discharge of patients at very low risk for myocardial infarction. We previously developed a single-sample rule-out pathway using the ARCHITECT hs-cTnI assay to risk stratify patients with suspected acute coronary syndrome. In a secondary analysis of the POC-ET (Point of Care Evaluation of High-sensitivity Cardiac Troponin) study, we evaluated performance of risk stratification with the Alinity hs-cTnI assay. Methods: Patients presenting with possible myocardial infarction in the POC-ET (NCT05665127) study were included. The primary outcome was type 1, 4b or 4c myocardial infarction or cardiac death at 30 days. Cardiac troponin I (cTnI) was measured in stored materials using the ARCHITECT and Alinity hs-cTnI assays. The sex-specific 99th percentile upper reference limit (URL) are 34 ng/L in men and 16 ng/L in women for both assays. Agreement was assessed with Bland-and-Altman limit of agreement method, Passing Bablok regression, and Pearson's correlation coefficient. Distributions of presentation measurements were compared with Kolmogorov-Smirnov test. Performance was evaluated in the overall population and prespecified subgroups. The negative predictive value (NPV) and sensitivity were determined and proportion of patients identified as low, intermediate, and high risk were calculated and modelled using ordinal logistic regression. Results: In 986 patients (60 [51-70] years, 38% female), 78 (7.9%) had a primary outcome. Strong agreement was found in the raw cTnI measurements (99% samples within the Bland-Altman limit of agreement; correlation coefficient r: 0.967 (95% CI 0.964-0.969, P<0.001); Passing Bablok regression: slope 1.12 [1.11-1.13], intercept -0.16 [-0.18 to -0.13]). At presentation, distributions of cTnI measurements by the two assays were similar (P=0.810). Both assays showed comparable diagnostic performance using a risk stratification threshold of <5 ng/L and the sex-specific diagnostic threshold, with the same NPV (Alinity 100 [99.7-100]% versus ARCHITECT 100 [99.7-100]%) and sensitivity (Alinity 100 [97.3-100]% versus ARCHITECT 100 [97.3-100]%). Similar proportions of patients stratified as low- (Alinity 67% versus ARCHITECT 67%), intermediate-risk (23% versus 24%) and high-risk (10% versus 9%) at presentation with minor reclassification. Similar efficacy was observed across subgroups stratified by sex, age, history of myocardial infarction, renal function, and symptom duration. Conclusions: The Alinity hs-cTnI and the ARCHITECT hs-cTnI assays can be used interchangeably in the assessment of suspected myocardial infarction with comparable safety and efficacy.
Da Costa, A.; Yvorel, C.; Romeyer, C.; Groussin, P.; Barengo, A.; Mohammed, R.; Azarnouch, K.; Grand, N.; Boukhris, M.; Benali, K.
Show abstract
Background. Durable mitral isthmus (MI) block remains challenging in persistent atrial fibrillation (PeAF) ablation. Recent epicardial vein of Marshall (VoM) recordings have shown incomplete MI transmurality and time-dependent conduction recovery after pulsed field ablation (PFA). Whether systematic VoM ethanol infusion (VoM-EI) followed by focal PFA provides stable acute MI block remains unknown. **Objectives.** To assess the incidence, timing, and procedural implications of early MI conduction recovery after systematic VoM-EI followed by focal Sphere-9 PFA. Methods.In this prospective single-center study, 55 consecutive patients undergoing first ablation for symptomatic PeAF with planned MI ablation were screened. VoM-EI was systematically attempted before left atrial access and successfully performed in 51 (92.7%), who constituted the study cohort. Pulmonary vein isolation, roof-line, and MI ablation were performed with the Sphere-9? lattice-tip catheter. After bidirectional MI block, conduction was systematically reassessed during a standardized 30-minute waiting period. Results.Mean age was 70.3 {+/-} 8.2 years, and 36 patients (70.6%) were men. Initial bidirectional MI block was achieved in 50/51 patients (98.0%). During the waiting period, conduction recovered in 9/50 (18.0%; 95% CI, 9.8%-30.8%), at a median of 16 minutes (IQR, 10-20; range, 8?23). Six of 9 patients with recovery (66.7%) required targeted coronary sinus (CS) ablation. Block was restored in all 9, yielding a final block rate of 50/51 (98.0%). Median procedure duration was 82 minutes (IQR, 73-95), with no major complications. Conclusions. Immediate bidirectional MI block was not synonymous with stable block. Despite systematic VoM-EI followed by focal Sphere-9 PFA, conduction recovered in approximately one in five patients, including beyond 20 minutes, and two thirds required targeted CS ablation. These findings support standardized 30-minute reassessment and targeted CS interrogation rather than reliance on immediate block. Chronic invasive remapping is required to determine whether this strategy improves long-term MI block durability.
Giordano, S.; Corcione, N.; Morello, A.; Cimmino, M.; Albanese, M.; Ferraro, P.; Vecchione, G.; Amat-Santos, I. J.; Giordano, A.; Biondi-Zoccai, G.
Show abstract
Background: Bailout cardiac surgery during transcatheter aortic valve replacement (TAVR) is uncommon but remains associated with substantial morbidity and mortality. Although registries have described its incidence and major causes, they often provide limited detail regarding device-related failure mechanisms, attempted transcatheter rescue, and the clinical pathway leading to surgical conversion. We aimed at analyzing post-marketing safety reports from the U.S. Food and Drug Administration (FDA) Manufacturer and User Facility Device Experience (MAUDE) database to characterize the mechanisms, management strategies, and reported outcomes of bailout surgery during or shortly after TAVR. Methods: We retrospectively analyzed FDA MAUDE reports received from July 1, 2016, through June 30, 2026. Eligible reports described unplanned urgent or emergent open cardiac surgery during or immediately after TAVR. Candidate reports were screened, adjudicated, and deduplicated at the clinical-event level. Events were classified by precipitating complication, transcatheter rescue, operative pathway, and reported outcome. Associations were evaluated using permutation tests, Fisher exact tests with Benjamini?Hochberg correction, adjusted regression models, and sensitivity analyses. Results: After screening 43,239 initial reports, we identified 376 bailout-surgery events, with survival status was documented in 254, including 104 deaths and 150 survivors, corresponding to 40.9% reported mortality. Valve embolization, migration, or malposition was the most frequent complication phenotype (32.4%), whereas ventricular perforation or laceration was associated with the highest mortality (74.1%; OR, 4.86; 95% CI, 1.97?11.99). Mortality differed across complication phenotypes (p<0.001) and operative pathways (p<0.001), but not across transcatheter rescue pathways (p=0.355). Valve explantation with SAVR was associated with lower reported mortality (18.9%; OR, 0.29; 95% CI, 0.12?0.69), whereas unspecified surgery or access/support alone was associated with higher mortality (56.9%; OR, 3.04; 95% CI, 1.80?5.12). Ancillary analyses identified potential platform-specific differences in complication and management patterns, while bailout timing was not independently associated with mortality after adjustment. Conclusions: In this MAUDE analysis, bailout cardiac surgery after TAVR was most commonly precipitated by valve embolization, migration, or malposition, whereas ventricular perforation or laceration was associated with the highest reported mortality. Outcomes differed across complication and operative pathways but not across transcatheter rescue strategies or bailout timing after adjustment. These findings identify clinically relevant post-marketing safety signals but should not be interpreted as incidence estimates, comparative device risks, or causal treatment effects.
Perlman, A.; Goldstein, N.; Goldman, M.; Shapiro, M.; Barash, E.; Bar, A.; Raveh, T.; Tordjman, E.; Schussheim, H.; Dormont, F.; Matalon, O.
Show abstract
Background. Cardiovascular-outcomes trials are lengthy, costly, and associated with substantial uncertainty prior to readout. In-silico trial simulation using real-world data (RWD) has emerged as a potential tool to support earlier decision-making; however, evidence of prospective predictive validity, generated prior to trial result disclosure, remains limited. Methods. We applied a semi-mechanistic machine learning framework integrating real-world patient data with biologically informed drug representations to prospectively simulate the VESALIUS-CV trial evaluating evolocumab versus placebo. The simulation model was trained on a combination of patient-level real-world data and a drug-centric knowledge graph and validated for both patient-level and trial-level retrospective predictive performance. The model was then used to simulate VESALIUS-CV before public disclosure of trial results, using a locked model and prespecified eligibility criteria and primary endpoint aligned with the clinical protocol. A patient-level time-to-event model was used to generate virtual trial arms, from which cumulative incidence curves, hazard ratios, confidence intervals, and p-values for major adverse cardiovascular events (MACE) were estimated. Results. In retrospective validation, the model demonstrated strong patient-level discrimination, with time-dependent ROC-AUC values ranging from 0.80 to 0.90 across follow-up horizons. For trial-level validation, 22 randomized cardiovascular-outcomes trials were simulated, and hazard ratios for 3-point MACE across 24 between-arm comparisons showed consistent directional agreement and quantitative correlation with published results such that the model accurately predicted trial success, achieving an F1 score of 0.83, with precision of 0.79 and sensitivity of 0.89. In a fully prospective application, the simulation predicted a statistically significant reduction in 3-point MACE with evolocumab versus placebo, estimating a hazard ratio of 0.78 (95% CI, 0.70-0.87) at 54 months. These predictions were consistent with the subsequently reported VESALIUS-CV results, which demonstrated a hazard ratio of 0.75 (95% CI, 0.65-0.86) at 55 months of median follow-up. Conclusions. In a fully prospective setting, a RWD-driven, AI-based simulation accurately predicted the direction, magnitude, and temporal dynamics of treatment effects observed in the VESALIUS-CV trial. These results demonstrate that in-silico trial simulation can anticipate clinical outcomes in the prospective setting, supporting its use as a complementary tool for early decision-making, trial design optimization, and de-risking in cardiovascular drug development.
Xiang, S.; He, H.; Xie, Z.; Cheng, C.-Y.; Li, H.; Liu, D.
Show abstract
Agentic workflows can coordinate modelling, but balancing predictive performance, measurement burden and reproducibility is unclear. We developed DXA Agent, an agentic workflow for dual-energy X-ray absorptiometry (DXA) outcomes integrating planning, feature-model refinement, tools, provenance and hypothesis-generating interpretation. Models were independently developed and tested in UK Biobank (5,318 participants) and the National Health and Nutrition Examination Survey (NHANES; 3,777 participants), using cost-efficient and no-limit strategies. Across 20 UK Biobank and three NHANES bone mineral density sites, cost-efficient models achieved lower RMSE and higher R2 than the best conventional comparator, with median relative RMSE reductions of 10.9% and 9.9%, respectively. Classification was task dependent: UK Biobank osteoporosis averaged AUROC 0.839 and PR-AUC 0.182, whereas NHANES performance was comparable with conventional models. Higher-burden features did not consistently improve prediction. These retrospective, cohort-internal findings position DXA Agent as an inspectable, measurement-burden-aware research workflow requiring independent prospective validation.
Mathew, Z.; Mehta, R.; Kim, S.; Jeyaraj, J.; Asif, T.
Show abstract
Background: Primary malignant cardiac tumors (PMCTs) are rare and histologically heterogeneous. Objective: To compare demographics, specific ICD-O-3 morphologies, first-course treatment patterns, annual registered case counts, and unadjusted overall survival between soft-tissue and hematologic PMCTs. Methods: We identified 730 PMCT cases diagnosed from 2000 to 2021 in SEER 18 (ICD-O-3 topography C38.0). Histologic lineage was assigned from ICD-O-3 morphology. Comparative analyses included soft-tissue (n=458) and hematologic (n=212) tumors. First-course variables were primary-site surgery, chemotherapy (yes versus no/unknown), and radiotherapy (radiation versus none/unknown). Groups were compared with chi-square tests. Overall survival was estimated with Kaplan-Meier methods; follow-up was truncated at 120 months. Results: Soft-tissue PMCTs occurred predominantly at ages 45-64 years (67.9%), whereas hematologic PMCTs occurred predominantly at age [≥]65 years (63.2%; p<0.001). Men comprised 59.9% of hematologic and 49.3% of soft-tissue cases (p=0.014). The leading soft-tissue morphology was hemangiosarcoma/angiosarcoma (ICD-O-3 9120/3; 201/458, 43.9%); synovial sarcoma accounted for 20/458 cases (4.4%). Diffuse large B-cell lymphoma, NOS, accounted for 131/212 hematologic tumors (61.8%). Any primary-site surgery was recorded in 66.6% of soft-tissue versus 15.6% of hematologic cases (p<0.001). Chemotherapy was recorded in 67.5% versus 51.1% (p<0.001), and radiotherapy in 9.0% versus 20.5% (p<0.001). In exploratory Kaplan-Meier analyses, hematologic patients with recorded chemotherapy had higher unadjusted 120-month overall survival than those without recorded chemotherapy (42.0% versus 12.2%; log-rank p=7.5x10-). Radiation-associated survival differences were not statistically significant in either lineage. Conclusions: Soft-tissue and hematologic PMCTs have distinct age distributions, named histologies, and first-course treatment patterns in SEER. These findings describe registry coding and do not establish treatment effectiveness or population incidence.
Choi, L.; McNeer, E.; Beck, C. A.; Neul, J. L.
Show abstract
Bayesian borrowing of external information can improve trial efficiency, particularly in pediatric and rare disease settings where patient populations are limited, but may introduce bias and inflate the Type~I error rate when the trial differs from external studies. Recent U.S. Food and Drug Administration (FDA) draft Bayesian guidance emphasizes careful evaluation of external information, prior specification, and assessment of operating characteristics. This paper compares three meta-analytic-predictive (MAP)-based methods for Bayesian borrowing: the MAP prior, robust MAP (RMAP) prior, and self-adapting mixture (SAM) prior. An adaptive platform trial design in Rett syndrome is used as a case study. Simulation studies evaluate frequentist operating characteristics under varying prior--data conflict, between-study heterogeneity, treatment effects, and clinically significant differences (CSDs) for the SAM prior. The MAP prior achieved the greatest efficiency when external and current data were compatible but exhibited the largest bias under substantial prior--data conflict. The RMAP priors improved robustness through fixed robust-component weights, whereas the SAM prior adaptively adjusted borrowing and was less sensitive to prior--data conflict while retaining efficiency gains when the data were compatible. Although the CSD influenced the degree of adaptive borrowing, as reflected by effective sample size, it had only a modest impact on frequentist operating characteristics. Sensitivity analyses using a skeptical robust component yielded similar qualitative conclusions, while accentuating the differences between the MAP and RMAP priors. These findings provide guidance for evaluating and selecting MAP-based borrowing strategies before trial implementation, particularly in rare disease settings, consistent with current FDA recommendations.
Muniz-Chicharro, A.; Tanriver, G.; Gora, A.
Show abstract
Summary: Prot2Surf is a software tool designed for the characterization and prediction of protein association to surfaces. In this application note, Prot2Surf was tested using catalytic domains of the lytic polysaccharide monooxygenases (LPMOs), interacting with native surfaces. The results show that the software can efficiently analyze key binding features, including protein-surface distances, distances between catalytically reactive atoms, and the orientation angle between surface chains and the protein. These features are essential for distinguishing productive binding poses in these protein-surface systems and for understanding interaction patterns that provide guidance on protein engineering. Prot2Surf performs these analyses within seconds to a few minutes, providing a fast and accessible framework to post-process and characterize protein-surface encounter complexes. Availability and implementation: Prot2Surf, which is written in Fortran90, is documented and freely available as open source on GitHub: https://github.com/TUNNELING-GROUP/Prot2Surf. In order to run Prot2Surf, users should also install the SDA software package which is freely available at https://www.h-its.org/downloads/sda7/.
Wojcik, S.; Rulkiewicz, A.; Domienik-Karłowicz, J.
Show abstract
Large language models perform well on medical examinations, but users routinely challenge their answers and invoke professional roles, and it is unclear what a system does when a medical credential and a stated task-specific accuracy point in opposite directions. In a factorial experiment on 480 items from four Polish specialty examination sets and three consumer large language model systems (ChatGPT, Claude, Gemini), each item and system received eleven independent conversations. Conditions crossed attributed source role (medical student, experienced specialist), stated prior accuracy on similar questions (2/10, 8/10) and suggestion correctness. The primary outcome was adoption of a prespecified incorrect option when the baseline answer matched the official key, comparing a specialist described as 2/10 with a student described as 8/10. Baseline agreement with the key was 87.2% across 15,683 analyzable conversations. The incorrect option was adopted more often from the specialist described as 2/10 than from the student described as 8/10 (10.2% vs. 7.6%; adjusted risk difference +2.82 percentage points, 95% CI +0.65 to +4.99). Estimates varied across the three systems and only one system-specific interval excluded zero. In a prespecified exploratory analysis with a shared eligibility rule, correct suggestions were adopted far more often than incorrect ones (risk difference +35.7 percentage points, 95% CI +30.8 to +40.7), indicating selective rather than indiscriminate compliance. An incorrect suggestion from a specialist with low stated accuracy was therefore slightly more influential than the same suggestion from a student with high stated accuracy, although the difference was modest and varied across systems. Agreement reached only after a user has disclosed a preferred answer should not automatically be treated as an independent second opinion, and medical large language model systems should be evaluated on how they revise answers after such disclosure, not solely on initial accuracy.
Otte, J. H.; Cartagena, A.
Show abstract
Background. A primary constraint on the capacity of EMS programs to meet industry demand is psychomotor instruction and verification, requiring direct observation of each student by a qualified evaluator. Whether AI video analysis can relieve it is untested; none has been applied to EMS skill examination or compared with human examiners. Objective. To quantify human EMS evaluator inter-rater reliability and evaluate an AI video-analysis platform against it. Methods. In a prospective, fully crossed study, five certified EMS evaluators and an AI platform independently scored identical video-recorded EMT performances of cervical collar application (n=15), bag-valve-mask (BVM) ventilation (n=14), and medical assessment (n=15) on dichotomous checklists with critical-failure criteria. Agreement was assessed at item, score, and decision levels using Fleiss' kappa, Krippendorff's alpha, Gwet's AC1, and ICC(2,1)/ICC(2,k). Results. Human item agreement was moderate (kappa 0.409 to 0.467), as was single-rater reliability (ICC(2,1) 0.539 to 0.694), against good panel reliability (ICC(2,k) 0.854 to 0.919). Recorded pass/fail agreement was fair (kappa 0.297 to 0.388) and critical-failure agreement near zero for two skills (kappa 0.028, 0.119). AI alignment tracked rubric observability rather than task complexity: r = 0.857 (collar, exceeding every human), -0.173 (BVM), 0.664 (medical), and it was most lenient on two skills. Conclusions. Human evaluators are an imperfect standard, especially on critical failures. The AI was a legitimate additional rater where checklist items were discrete and visually verifiable, but not where credit required judging continuous quantities such as ventilation rate, volume, or suction duration. Defensible uses are formative and archival, not summative. These results reflect an early, non-specialist configuration: a baseline, not a limit.